Skip to content

Kimik3 madengine integration - #242

Open
Cemberk wants to merge 94 commits into
ROCm:developfrom
Cemberk:kimik3-madengine-integration
Open

Cemberk wants to merge 94 commits into
ROCm:developfrom
Cemberk:kimik3-madengine-integration

Conversation

@Cemberk

@Cemberk Cemberk commented Sep 1, 2026 •

Copy link
Copy Markdown

(built on top of existing kimi k3 pr for madengine enablement and testing)

Summary

Adds Kimi-K3 multinode inference to MAD and makes the multinode vLLM / SGLang workloads run, fail and report the same way whether they are launched through madengine or submitted directly with sbatch. It also adds a docs/ folder that takes a new reader from zero to running, configuring and extending these workloads.

Start here: docs/README.md → docs/multinode-running.md.

Depends on: ROCm/madengine#203. It is needed for exact Dockerfile matching, for counting a failed SLURM deployment as a failure, and for the slurm_multi collector reading a card's multiple_results.

What's in this PR

Kimi-K3 on MI300X / MI355X: 7 new cards, 1 image

Card Launcher Shape
pyt_vllm_kimi-k3_mi300x_wideep_moriep scripts/vllm_multinode (colocated) 2 nodes, wide EP, MoRI-EP
pyt_vllm_kimi-k3_mi300x_wideep_allgather colocated 2 nodes, wide EP, all-gather
pyt_vllm_kimi-k3_mi300x_pp2xtp8 colocated 2 nodes, PP2 × TP8
pyt_vllm_kimi-k3_mi300x_pp2xtp8_way4 colocated same workload, expressed through mad-config.kimi-k3.yaml
pyt_vllm_disagg_mori_kimi-k3-mxfp4 scripts/vllm_dissag (disaggregated) 2P/2D, TP2 × DP8, 4 nodes
pyt_vllm_disagg_mori_kimi-k3 disaggregated 2P/2D, TP2 × DP8, 4 nodes
pyt_vllm_disagg_mori_kimi-k3-mxfp4_mi355x disaggregated MI355X, TP1 × DP16 (bring-up)
  • One image for all seven cards: docker/vllm_kimi_k3.ubuntu.amd.Dockerfile.
    • It builds for MAD_SYSTEM_GPU_ARCHITECTURE (gfx942 or gfx950, no default).
    • AITER is built from source at the Kimi-K3 release commit (ROCm/aiter 68e42f5f), with flydsl 0.2.4. Nothing is copied from a second image.
  • Weights: REQUIRE_LOCAL_WEIGHTS=1 on the Kimi cards requires local NVMe copies. MODEL_WEIGHTS_NAME selects the checkpoint directory when it is staged under another name (for example Kimi-K3 vs Kimi-K3-MXFP4).

Launchers: every failure ends the job with the reason in the log

  • A server that dies at start-up, including one worker running out of GPU memory, fails every node fast with its first error lines. The job no longer waits out a timeout.
  • The vLLM and SGLang disagg launchers fail the job when a node's container failed, or when the benchmark produced no perf.csv.
  • The rixl/DeepEP connector honours the recipe's AITER, KV and cudagraph settings instead of overriding them.
  • Benchmark clients request the model under the recipe's --served-model-name.

Results: perf.csv says what happened, cell by cell

  • Both parsers (vLLM and SGLang) read the log one cell at a time. A cell that stalled, printed no result, lost any request, or measured zero throughput is a FAILURE row. It is no longer a missing row or a false SUCCESS.
  • For NIAH, a context length with any errored request, or NO-RESULT, is a FAILURE row.
  • long_context runs now write perf.csv.

Configuration: one source for each setting

  • Allocation defaults are madengine's SLURM presets (partition, GPUs per node, exclusive).
  • Site facts (weights, fabric, ports, logs) live in scripts/common/cluster.sh.
  • Model tuning lives in each launcher's models.yaml recipe.
  • There are no per-site JSON files. A run passes only what differs, in --additional-context.
  • docs/configuration.md lists every layer, what wins, and where to change what.

How to run

Install madengine with the changes this PR depends on:

pip install git+https://github.com/Cemberk/madengine.git@rocenvflag

Through madengine

From the repo root on a SLURM login node. The Kimi cards carry DOCKER_IMAGE_NAME: "<supply-your-image>", so either build and push an image or point at one.

# Build and push the image for the compute nodes' GPU, then run.
# A login node has no GPU, so name the arch.
madengine build --tags pyt_vllm_disagg_mori_kimi-k3-mxfp4 --registry docker.io/<namespace> \
    --additional-context '{"slurm": {"time": "06:00:00"},
                           "docker_build_arg": {"MAD_SYSTEM_GPU_ARCHITECTURE": "gfx942"}}' \
    --manifest-output build_manifest.json
madengine run --manifest-file build_manifest.json --timeout 21600 --live-output -o perf.csv

# Or use an image you already have:
madengine build --tags pyt_vllm_kimi-k3_mi300x_wideep_moriep --use-image <registry>/<repo>:<tag> \
    --additional-context '{"slurm": {"time": "04:00:00"}}'
madengine run --manifest-file build_manifest.json --live-output
  • What you usually need to pass: only the time limit, because the cards say 24:00:00 and most partitions cap lower.

  • Where the rest comes from: the node count is in each card. The partition, GPUs per node and exclusivity are madengine's presets. Model and log paths default in scripts/common/cluster.sh.

  • On a cluster that differs, add only the keys that differ:

    {"slurm": {"partition": "<partition>", "time": "06:00:00"},
     "env_vars": {"MODEL_DIR": "/path/to/models", "LOG_PATH": "/path/to/shared/logs"}}
  • If a node stages the checkpoint under another name, add "MODEL_WEIGHTS_NAME": "Kimi-K3" under env_vars.

Standalone with sbatch

Each launcher takes the card's env_vars as its contract. Export them, name the image, and submit with the allocation. sbatch options go before the script.

# Disaggregated vLLM (DeepSeek-V3, MoRI-IO + MoRI-EP, 1P/1D)
cd scripts/vllm_dissag
export DOCKER_IMAGE_NAME=<registry>/<image>:<tag> \
       MODEL_NAME=DeepSeek-V3 xP=1 yD=1 WIDE_EP=1 RUN_MORI=1 RUN_DEEPEP=0 \
       BENCHMARK_COMBINATIONS=1024/1024
sbatch --partition=amd-rccl --nodes=2 --ntasks=2 --gpus-per-node=8 --exclusive \
       --time=06:00:00 --export=ALL run_xPyD_models.slurm

# Colocated vLLM (Kimi-K3 PP2 x TP8, NIAH)
cd scripts/vllm_multinode
export DOCKER_IMAGE_NAME=<registry>/<image>:<tag> MODEL_NAME=Kimi-K3 \
       TP_SIZE=8 PP_SIZE=2 ENABLE_EP=0 AITER_SITUV2_A8W4=1 GPU_ARCHS=gfx942 \
       REQUIRE_LOCAL_WEIGHTS=1 BENCHMARK_SCRIPT=niah NIAH_WORDS=10000,50000,100000,200000
sbatch --partition=<partition> --nodes=2 --time=06:00:00 --export=ALL run_multinode.slurm
  • A run is reported identically either way:
    • the job exits non-zero on failure;
    • results land in $LOG_PATH/<job>/perf.csv, one row per cell with SUCCESS or FAILURE;
    • every node's server log is next to them.
  • More detail: see docs/multinode-running.md (inside an salloc, weights, fabric, logs, fail-fast) and docs/kimi-k3.md.

Checks that need no GPUs

bash scripts/vllm_dissag/tests/argv_assert.sh          # the vllm serve argv/env each connector and recipe builds
bash scripts/vllm_dissag/tests/parse_to_csv_assert.sh  # vLLM sweep, NIAH and long_context result rows
bash scripts/sglang_disagg/tests/parse_to_csv_assert.sh

Validation status (MI300X, gfx942)

Workload Status
vLLM disagg: NIXL Llama-3.3-70B-FP8, MoRI DeepSeek-V3, MoRI GLM-5.1-FP8 Passing, every sweep cell 8 → 512
SGLang disagg: MoRI-IO Qwen3-32B, MoRI-DP DeepSeek-V3, Mooncake Qwen3-32B / DeepSeek-V3, agentic Qwen3-32B Passing
Kimi-K3 colocated: wide EP MoRI-EP, wide EP all-gather, PP2 × TP8 (both forms) Passing, NIAH 10/10 at 10k–200k words
Kimi-K3-MXFP4 disagg 2P/2D Weights resolve and both servers start on 4 nodes; end-to-end NIAH run in progress

Known issues

  • DeepEP DeepSeek-V3 (pyt_vllm_disagg_deepep_deepseek-v3, pyt_vllm_disagg_nixl_deepseek-v3):
    • start-up now works;
    • prefill fails at concurrency ≥ 128 with DeepEP error: CPU recv timeout in high-throughput intranode dispatch;
    • needs a DeepEP-side fix. The results mark those cells FAILURE.
  • pyt_vllm_disagg_mori_kimi-k3 (the Kimi-K3 recipe): its memory budget (40e9 KV cache and a 16 GiB MoRI heap on top of ~148.6 GiB of TP2 weights) does not fit a 192 GB MI300X. It is left unchanged pending the recipe owner. pyt_vllm_disagg_mori_kimi-k3-mxfp4 is the validated Kimi disagg recipe.
  • SiTUv2 warning on the gfx942 int4 MoE path. The Kimi image's vLLM warns that its AITER build "ignores the SiTUv2 activation on the packed-int4 MoE path".
  • GLM-5.1-FP8 and Mooncake DeepSeek-V3 show lower throughput at concurrency 512 than at 256, with no failed requests.

Review guide

  • docker/vllm_kimi_k3.ubuntu.amd.Dockerfile: the single Kimi image.
  • scripts/vllm_dissag/, scripts/sglang_disagg/, scripts/vllm_multinode/: launchers, recipes (models.yaml), cards (models.json), parsers.
  • scripts/common/cluster.sh: site facts shared by all three launchers.
  • docs/: new. The existing READMEs are unchanged except where they now point into docs/.

Cemberk and others added 14 commits August 13, 2026 17:17
… selector

The xPyD harness already models pools that span multiple nodes (xP/yD count
NODES, roles are MASTER/CHILD via --data-parallel-start-rank + --headless), but
the wideEP path hardcoded TP=1: dp_size was xP*GPUS_PER_NODE and the connector
emitted a literal `-tp 1`. That makes any model whose REPLICATED (non-expert)
weights exceed one GPU inexpressible - e.g. Kimi-K3 on MI300X, where TP1/DP16
needs 190.7 GiB/GPU (> 192 GB HBM) and TP2 is required to fit.

- vllm_disagg.sh: topology math is TP-aware. dp_per_node = GPUS_PER_NODE/TP_SIZE
  feeds dp_size, dp_size_local and both start-rank calculations. TP_SIZE defaults
  to 1 and is validated to divide GPUS_PER_NODE.
- moriio.sh: emit -tp ${TP_SIZE}; clamp --api-server-count to dp_size (it must be
  <= data-parallel-size or the frontend DP balancer routes to ranks that do not
  exist - only reachable once TP>1 shrinks dp_size below GPUS_PER_NODE).
- moriio.sh: advertise the peer pool's node IPs as moriio_pod_hosts when a pool
  spans >1 node. Without it the connector falls back to the peer MASTER only and
  KV writes aimed at ranks on a peer CHILD node silently miss, so those ranks
  decode with no context. Emitted only when xP>1 || yD>1.
- run_xPyD_models.slurm: forward TP_SIZE/GPUS_PER_NODE, wire BENCHMARK_SCRIPT=niah
  to the already-present benchmark_niah.sh, forward NIAH_* knobs.

Default-neutral: at xP=1 yD=1 TP_SIZE=1 - the topology every registered entry
uses - DRY_RUN output is byte-identical to develop across all four node roles.
At xP=2 yD=2 TP_SIZE=2 it emits tp=2, dp_size=8, dp_local=4, child start-rank 4
(= EP16 per pool over 16 GPUs).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…Kimi-K3 recipe

Two parts.

1) moriio.sh wideEP dropped ${model_args[@]} on the floor. models.yaml documents
   per-model dp: tuning as supported ("both connectors now append it") and rixl's
   deepep path does append it, but the moriio wideEP branch parsed the flags,
   logged them, and then emitted an argv without them. Verified with a sentinel
   flag: present in the log line, absent from the command. Now appended, matching
   rixl. Every existing wideEP model has an empty dp: block, so emitted argv is
   unchanged for all of them.

2) Kimi-K3 recipe, entirely as data: env: block for the gfx942 knobs
   (VLLM_ROCM_USE_AITER_MLA=0 - the AITER MLA kernel is gfx950-only -
   AITER_SITUV2_A8W4=1 for the packed-int4 SiTUv2 MoE path, KDA conv state layout,
   40e9 KV cache for >600K contexts, per-role cudagraph + MoRI backends), and a
   dp: block for the serve flags the connector does not emit (reasoning parser,
   1M max-model-len, batched-tokens pinned at 2048, MoE quantization-config).
   Registered in MORI_EP_VALID_MODELS and WIDE_EP_ONLY_MODELS.

   The JSON quantization-config carries its own shell quotes: the launcher
   word-splits dp: via eval, which would otherwise strip the JSON's double quotes
   and hand vLLM invalid JSON. Verified the emitted value parses as JSON.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docker/pyt_vllm_kimi_k3_mi300x.ubuntu.amd.Dockerfile
  The shared vllm_disagg_inference stack with the pins K3 needs on gfx942
  (MoRI 1.2.2, AITER 0.1.19 + flydsl 0.2.4, the K3+MoRIIO vLLM commit) plus the
  K3-aware AITER graft from the public vendor image - without that graft the K3
  MoE profiling shape finds no tuned FlyDSL config, falls back to a heuristic
  kernel and aborts LLVM inside determine_available_memory.

  Differences from the recipe this is ported from:
  - every source pinned to an immutable SHA, not a fork branch name (branches on
    personal forks can be force-pushed; MAD needs the image rebuildable later)
  - GH_TOKEN build-arg dropped: all three repos are public, and a token passed
    this way is recorded in image metadata
  - the flydsl 0.2.4 re-pin happens at build time, so the launcher no longer runs
    pip install inside every container at serve time
  - WITH_NIXL defaults to 0 (K3 uses the moriio connector only)
  - no runtime patchers: all connector/KDA fixes are committed in the pinned vLLM

scripts/vllm_multinode/
  Colocated (single-instance) multi-node serving - the counterpart to
  vllm_dissag, for models that do not fit one node but want lowest single-request
  latency rather than disagg throughput. TP within a node, PP across nodes.

  It owns only node discovery, the container launch and the head/worker split;
  models.yaml, socket_barrier.py, benchmark_xPyD.sh, benchmark_niah.sh and
  parse_to_csv.py are reused from vllm_dissag (scripts/ is mounted whole), so the
  model recipe and the CSV pipeline have one home.

  The three K3 colocated variants collapse to this one script: dry-run confirms
  pp2xtp8 / allgather / moriep argv differ only by --enable-expert-parallel and
  --all2all-backend, so they become three models.json entries rather than three
  near-identical run.sh copies.

Not built here - this environment has no docker and one 8-GPU node; the harness
is verified by DRY_RUN argv inspection only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
models.json: four entries, all skip_gpu_arch gfx950 (the existing four Kimi-K3
entries are gfx950/TP8 single-node and all carry skip_gpu_arch gfx942, so K3 is
currently skipped entirely on MI300X - this is that gap):

  pyt_vllm_disagg_mori_kimi-k3               -N 4  xP2 yD2 TP2 -> EP16/pool
  pyt_vllm_kimi-k3_mi300x_pp2xtp8            -N 2  TP8 x PP2, no EP
  pyt_vllm_kimi-k3_mi300x_wideep_allgather   -N 2  + EP, allgather_reducescatter
  pyt_vllm_kimi-k3_mi300x_wideep_moriep      -N 2  + EP, mori_low_latency

The three colocated entries differ only in ENABLE_EP / ALL2ALL_BACKEND /
AITER_SITUV2_A8W4, which is why they share one launcher.

Result reporting: BENCHMARK_SCRIPT=niah previously produced logs and nothing
else, so a model declaring multiple_results would have reported no metric.
- parse_to_csv.py gains --niah: parses benchmark_niah.py output and writes the
  same 29-column madengine perf.csv as the throughput path (verified identical
  column sets). One row per context size, performance = needles found /10. A
  size that errored is written as a FAILURE row with performance 0 rather than
  dropped, so a pass->crash regression is visible instead of silent.
- benchmark_niah.sh calls it.
- Both slurm launchers copy /run_logs/$SLURM_JOB_ID/perf.csv to
  ./perf_$MODEL_NAME.csv on completion: madengine resolves multiple_results
  against the job directory it launched from, not the container log mount.
  No-op for entries that do not declare multiple_results.

Verified: madengine discover goes 133 -> 137 with exactly these four added and
none removed; every referenced dockerfile and script exists; -N matches each
topology; the int4 quantization-config survives shell word-splitting as valid
JSON; DeepSeek moriio and rixl/deepep DRY_RUN output is byte-identical to
develop.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Ports the writeups from PR ROCm#193 next to the existing gfx950 K3 doc, keyed to the
madengine tags rather than to shell scripts: topology rationale (why MI300X needs
multi-node and why the disagg pool needs TP2), the gfx942 knobs and where each
one lives now, the NIAH results, and the three root-cause fixes.

Corrections carried in:
- the NIAH ceiling is 900K throughout. The upstream READMEs still said 300K in
  seven places after two later commits raised it to 500K and then 900K.
- states that the NIAH harness sizes context in WORDS (~1.3x tokens) while the
  results table is in tokens, so the two columns are not comparable.

benchmark/kimi_k3/README.md gains a pointer: all four entries there carry
skip_gpu_arch gfx942, which is exactly the gap the MI300X page fills.

Every tag, node count and knob value in the new page is checked against
models.json and models.yaml.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…metadata

The four MI300X Kimi-K3 entries could not run correctly on a cluster. Three
separate defects, all from the colocated launcher inheriting conventions that
only hold for the disaggregated one.

Node sizing. The entries carried `"args": "-N 4 -n 4"` on the assumption that
madengine forwards args to sbatch. It does not — args are appended to
`bash <model>.slurm`, and neither launcher parses $@, so the flags were dropped
and the job was submitted with the `slurm.nodes` default of 1. Sizing now uses
`distributed.nnodes`, which madengine reconciles into `#SBATCH --nodes`, and the
inert args are removed. `slurm.nodes` is deliberately left unset: setting it
selects the multi-node preset, whose `NCCL_SOCKET_IFNAME=eth0` is forwarded into
the container by run_xPyD_models.slurm and would override the fabric interface
on an RDMA cluster. `slurm.time` is set explicitly instead, since the
single-node preset's 12 h default is short for a full NIAH sweep.

NIAH model tag. benchmark_niah.sh requested MODEL_PATH as the model name,
correct for the disagg path where vLLM defaults served_model_name to the serve
argument. serve_colocated.sh passes `--served-model-name "$MODEL_NAME"`, so
every request 404'd and all three colocated entries — which default to
BENCHMARK_SCRIPT=niah — would have recorded a full sweep of FAILURE rows. The
harness now honors NIAH_MODEL when a launcher sets it, and the colocated
launcher sets it to the tag it actually serves.

Colocated run metadata. parse_to_csv derived nnodes, n_gpus, deployment_type and
tags from xP/yD. serve_colocated.sh exports xP=1 yD=0 only so the shared log
filenames stay unique, so a 2-node 16-GPU colocated run reported itself as a
1-node 8-GPU `disagg_1P0D` with a nixl backend it never used. Node count now
comes from NNODES (exported by both launchers), and a launcher whose shape is
not "xP prefill + yD decode" states its own identity via PERF_DEPLOYMENT_TYPE /
PERF_TAGS. The disagg path is byte-identical to before.

Also corrects the README, which documented the args-to-sbatch behavior that does
not exist and a quick start that cannot work now that the "<supply-your-image>"
placeholder is rejected rather than pulled.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
NUM_NODES is derived from the requested topology (xP+yD), not from the
allocation, and the node list was then truncated with `head -n $NUM_NODES`. A
short allocation therefore produced a silently wrong topology: pools were built
from nodes that were never allocated, and the run failed much later as a
connector handshake timeout with no indication of the real cause.

This mattered little while these models were always launched from a matching
`salloc`, but the sbatch path now sizes the allocation from the model card, so a
mismatch between card and cluster is a reachable state. Check up front and name
the three ways to fix it.

The colocated launcher already fails on the equivalent mismatch via its
TP*PP == NNODES*GPUS check, so it needs no counterpart.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ntract

parse_to_csv.py hand-wrote a full 29-column perf.csv, assembling node counts, GPU
counts, launcher, image and tags from the environment. Most of that is not the
workload's to know: madengine already owns it, and the guesses were wrong for the
colocated launcher, which sets xP=1 yD=0 only to keep log filenames unique.

The NIAH writer now emits a narrow CSV — model, performance, metric, status, plus
descriptive columns — and madengine merges in the run metadata via the entries'
`multiple_results` declaration. The columns that genuinely describe the workload's
own configuration rather than its placement (tp, pp, ep_backend, prefill_decode)
move to descriptive columns, matching how scripts/vllm/run_vllm.py already reports
tp/dtype/bs on the templated path. Nothing is lost; it is relocated to whichever
side actually knows it.

`status` stays explicit: an errored context size scores 0, and deriving status from
performance would file that real failure as a SUCCESS.

Scoped to leave every other workload alone. The NIAH writer is reachable only by the
four Kimi entries, which all declare `multiple_results`. The throughput writer keeps
its full-schema output by default, because the twelve disagg cards that use it
declare no `multiple_results` and madengine reads their CSV directly with no
metadata to merge — a narrow CSV there would drop every descriptive column. Their
output is byte-identical; `--narrow` is the opt-in for migrating one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…uted.nnodes

madengine reads only slurm.nodes when emitting `#SBATCH --nodes`
(deployment/slurm.py: `self.nodes = slurm_config.get("nodes", 1)`), and
nothing maps distributed.nnodes across — build_orchestrator copies nnodes
into deployment_config.distributed for launcher detection only. The four
MI300X model cards carried nnodes but no slurm.nodes, so every one of them
would have been submitted as a 1-node job. The launcher's
TP*PP == NNODES*GPUS_PER_NODE assertion catches it, but only after the
allocation is granted.

Add nodes/gpus_per_node to the four slurm blocks, and ship the site configs
that carry the values a model card cannot:

  - results_dir, absent from the key list madengine copies out of models.json,
    and required because the slurm_multi collector globs results_dir for
    perf*.csv rather than resolving multiple_results
  - MODEL_DIR / LOG_PATH, which otherwise default to /shared_inference

*.json is gitignored repo-wide, so the templates need explicit negations or
they never reach a clone.

Correct the two README claims that did not match madengine's behavior
(nnodes-sizes-the-allocation, multiple_results-is-resolved-first), document
that --account/--qos are unwired for slurm_multi and need SBATCH_ACCOUNT,
and drop --keep-model-dir from the SLURM examples — it is a local-Docker
flag that madengine ignores with a warning on SLURM.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Four defects, each found by running the recipes end to end on a real
MI300X SLURM cluster for the first time. None of them could have survived
a single successful run, which is consistent with these cards never having
been executed through madengine.

1. Build was not reproducible. Every source is pinned to a commit SHA, but
   pip resolves build dependencies in an isolated environment from PyPI at
   build time, so nothing pinned the toolchain. setuptools >= 80 added
       assert isinstance(self.compiler, CCompiler)
   to distutils' build_ext.build_extension, which MoRI's legacy
   Cython.Distutils.build_ext path violates, so the amd_mori wheel failed
   with "AssertionError: run() must precede build_extension()" while every
   pinned SHA was still correct (job 223849). PIP_CONSTRAINT is the only
   mechanism that reaches inside pip's build isolation -- installing
   setuptools in the image does not. Set globally so the AITER, vLLM and
   router stages cannot regress the same way. Build then completed in
   19 minutes, all 18 stages (job 223872).

2. RDMA library mounts used -e, which is true for directories. Docker
   creates a DIRECTORY at a bind-mount source that does not exist, so the
   first run on a node lacking e.g. libionic.so.1 leaves an empty directory
   behind, and every later run on that node mounts a directory over a file
   inside the image:
       OCI runtime create failed: ... not a directory
   Self-propagating and per-node, so it presents as intermittent. The
   surviving node then waits at socket_barrier.py forever for a peer that
   already died -- 4h10m of "Waiting for nodes. . ." before the wall clock
   (job 223909). Use -f so only regular files are mounted.

3. pyt_vllm_kimi-k3_mi300x_pp2xtp8 omitted the gfx942 MoE requirement that
   the two wideep variants carry. gfx942 has no scaled-MXFP4 MFMA and the
   a16w4 SiTUv2 heuristic FlyDSL kernel cannot codegen there, so the MoE
   must be requantized to packed int4. Without it the worker aborts inside
   determine_available_memory with
       LLVM ERROR: Do not know how to expand this operator's operand!
   and quantization_config=None in the engine config (job 224132). This is
   a hardware requirement, not an EP-specific tuning, so all three
   colocated cards now carry AITER_SITUV2_A8W4=1 and the int4 config.

4. A run that produced no results exited 0. The launchers warned about a
   missing perf CSV and returned success anyway, so madengine recorded
   exit_code=0 in its completion marker and the collector reported
   "0 successful, 0 failed" -- a hard engine crash was indistinguishable
   from a clean run. For a benchmark repo that is the worst possible
   reporting outcome. Both launchers now exit 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Each of these cost hours to diagnose on a real cluster, and in every case
the symptom is a long way from the cause -- an intermittent multi-node
hang that is actually a poisoned bind-mount path, a wheel build that fails
while every pinned SHA is correct, an LLVM abort that is a missing
quantization flag. Record them so the next person reads instead of
rediscovers.

Also record the cluster facts a model card structurally cannot carry: the
SLURM association wall-time cap (which silently parks a job as PENDING
rather than failing), account requirements, docker on compute nodes, and
the fact that a registry-less cluster has no supported path to distribute
an image -- --build-on-compute requires --registry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds Kimi-K3-MXFP4 (MI300X gfx942 MXFP4 MoE) to the vllm_dissag framework:
- models.yaml: K3 recipe env anchor + model entry (TP2×DP8, MoRI-EP)
- run_xPyD_models.slurm: allowlists + 2P/2D validation + JIT cache split
- moriio.sh: K3-gated topology (pod_hosts, api-server-count, headless KV)
- vllm_disagg.sh: pod hosts export + _model_config_to_array helper
- argv_assert.sh: K3 offline tests

Config-only integration; image build documented in README.

Co-Authored-By: Claude <noreply@anthropic.com>
…lsifiable

Both found on the first end-to-end MI300X run (job 224200), which served
Kimi-K3 and produced results -- but results that could not be trusted, from
a job that then would not exit.

Deadlock on teardown. serve_colocated.sh has the head kill only its own
server and exit, while every worker sits in `wait` on a headless vLLM that
nothing stops. srun waits on those tasks, so the job holds all nodes until
the wall clock: 4.5 hours of two idle exclusive MI300X nodes AFTER the
benchmark had already written perf.csv. The head now raises a sentinel on
the shared log volume and workers watch for it. The sentinel is raised from
an EXIT trap so a head that fails or times out cannot strand workers either.

NIAH could not distinguish truncation from retrieval failure. Scoring reads
content + reasoning_content, and K3 is a reasoning model served with
--reasoning-parser kimi_k3, so a trace that exhausts max_tokens is cut off
and only the earliest needles survive. At the old 2048 default, two of four
context sizes scored 1/10 and both listed exactly ANIMALS[0] -- the needle
placed first in the haystack. That signature is a truncated trace, not a
model failure, but finish_reason was never recorded so the two were
indistinguishable, and every row was written SUCCESS regardless.

  - record finish_reason and print it as finish=<reason>
  - raise the default max_tokens to 8192
  - keep the needle count as the metric, always reported, never dropped
  - mark a truncated row FAILURE and annotate the metric, so a measurement
    artifact cannot enter the results as if it were a model result

9/10 stays SUCCESS: the README documents ~9/10 at >=20K as a known RDMA
residual, so a strict 10/10 gate would flag known-acceptable behaviour.
Logs predating the finish= field still parse.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 1, 2026 19:37

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds Kimi-K3 MI300X support by integrating a new colocated multi-node vLLM launcher alongside updates to the existing disaggregated (prefill/decode) vLLM harness, plus reporting/benchmarking and documentation updates to make results consumable by madengine.

Changes:

  • Introduces a 2-node colocated vLLM SLURM+container launcher (vllm_multinode) with shared-model recipe support and shutdown coordination.
  • Extends the disaggregated launcher path to support TP_SIZE>1, multi-node pool host advertisement for MoRIIO, and NIAH benchmarking/reporting.
  • Adds Kimi-K3 model recipes/cards, MI300X-specific docs/site-config templates, and a pinned Dockerfile for the K3 MI300X image.

Reviewed changes

Copilot reviewed 15 out of 16 changed files in this pull request and generated 3 comments.

Show a summary per file
File Description
scripts/vllm_multinode/serve_colocated.sh Per-node in-container entrypoint for colocated multi-node vLLM serving + benchmarking.
scripts/vllm_multinode/run_multinode.slurm SLURM launcher that discovers nodes/IPs, runs per-node containers, and publishes perf CSV.
scripts/vllm_dissag/vllm_disagg.sh Adds TP-aware topology math and peer-pool host export for KV connector.
scripts/vllm_dissag/run_xPyD_models.slurm Adds Kimi-K3 support, TP_SIZE plumbing, NIAH option, safer RDMA mounts, perf CSV publish.
scripts/vllm_dissag/parse_to_csv.py Adds narrow-schema perf CSV mode, NNODES-aware metadata, and NIAH parsing/output.
scripts/vllm_dissag/models.yaml Adds Kimi-K3 MI300X serving recipe env + flags (MXFP4/int4 MoE path).
scripts/vllm_dissag/connectors/moriio.sh Adds multi-node pod host list to KV config; TP-aware api-server-count and tp flag.
scripts/vllm_dissag/benchmark_niah.sh Uses NIAH_MODEL when set; emits perf.csv for NIAH via parse_to_csv.py.
scripts/vllm_dissag/benchmark_niah.py Raises default max_tokens, records finish_reason, and prints parseable result lines.
models.json Adds Kimi-K3 MI300X disagg + colocated model-card entries with multiple_results CSV.
docker/pyt_vllm_kimi_k3_mi300x.ubuntu.amd.Dockerfile New pinned build for Kimi-K3 MI300X (MoRI/AITER/vLLM/router + graft step).
benchmark/kimi_k3/README.md Notes gfx942 requires multi-node sharding and points to MI300X recipes.
benchmark/kimi_k3/mi300x/slurm-config.disagg.json Site-config template for 4-node disaggregated MI300X runs.
benchmark/kimi_k3/mi300x/slurm-config.colocated.json Site-config template for 2-node colocated MI300X runs.
benchmark/kimi_k3/mi300x/README.md Full MI300X recipe documentation, tradeoffs, and madengine integration notes.
.gitignore Ensures MI300X site-config templates are not ignored.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +105 to +108
if [[ -n "${COLOCATED_EXTRA_ARGS:-}" ]]; then
eval "extra_args=(${COLOCATED_EXTRA_ARGS})"
serve_args+=("${extra_args[@]}")
fi
Comment on lines +149 to +151
docker rm -f $DOCKER_CONT_NAME 2>/dev/null || true
fuser -k 2223/tcp 2>/dev/null || true
sleep 2
Comment thread scripts/vllm_multinode/serve_colocated.sh Outdated
Job 239755 (Kimi-K3, 2x8 on oci-64) died in vLLM's startup barrier. Node 0
finished loading weights at 23:21:41 and entered in_the_same_node_as ->
torch.distributed.barrier; node 1 was still constructing MoE layers ~50 min
later. gloo gave up at 23:51:41, exactly 1800s in, and the master's TCPStore
died with it, so node 1 then failed with "Broken pipe" to :29500. The barrier
is all-or-nothing across all 16 ranks, so what has to stay small is the GAP
between nodes, not the absolute load time.

Five problems, addressed here:

1. --distributed-timeout-seconds does not reach the barrier that failed. It
   feeds ParallelConfig.distributed_timeout_seconds, which only configures the
   device/NCCL groups. The startup barrier runs on the gloo CPU group, built by
   GroupCoordinator from the separate cpu_distributed_timeout_seconds field.
   Unset, it falls back to PyTorch's stock 1800s -- which is exactly the
   deadline in the traceback, despite 7200 being passed. Pass
   --cpu-distributed-timeout-seconds too.

2. Cold AITER JIT compile on every rank, every run. The image points
   AITER_JIT_DIR/TRITON_CACHE_DIR/VLLM_CACHE_ROOT/COMGR_CACHE_DIR at
   /opt/vllm_cache, but this launcher mounted only /tmp/vllm_cache, a path
   nothing reads, inside --rm containers. Mount a host-persistent cache at
   /opt/vllm_cache keyed by image ID, as vllm_dissag/run_xPyD_models.slurm
   already does. Per-node cold compile is precisely the skew the barrier
   cannot absorb.

3. WORKER_PID was tee's PID. After `a | b &` the shell reports b, so every
   kill in this script hit tee and left vLLM running, and no liveness check
   was possible. Launch through process substitution instead.

4. A dead engine was indistinguishable from a slow one. The head polled the
   log for the full 4000s after its writer had already died, then reported a
   timeout rather than the traceback. Both wait loops now check liveness --
   via a helper that also rejects zombies, since an exited-but-unreaped child
   still answers kill -0. This also fixes a latent hang where a worker would
   spin forever on a dead engine.

5. Container env was a closed allowlist. Tuning any of the above from CI meant
   a new -e line and another MAD PR. COLOCATED_FORWARD_ENV names the variables
   to forward, and they are passed via docker --env-file so values containing
   spaces survive. The file is node-local, mode 600, and unlike the wrapper
   scripts is never archived as a build artifact.

Also adds an opt-out checkpoint page-cache pre-warm before the container
barrier (PREWARM_CHECKPOINT=0 to skip). Its per-node duration is the real
diagnostic: a large spread means the storage path is the problem and no
timeout will fix it.

Verified: bash -n clean on both files, and on the single-quoted srun body
extracted in isolation; env-file forwarding preserves spaces and skips unset
vars at mode 600; the liveness helper correctly reports an unreaped child as
dead; DRY_RUN=1 emits the new flag and honours an override, with
COLOCATED_EXTRA_ARGS still appended last.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings September 3, 2026 14:39

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The new run_multinode.slurm launcher has verified set -u-triggered failure paths (optional env expansion and hardcoded barrier port) that can prevent runs from starting or make barrier cleanup inconsistent.

Once you've addressed the issues Copilot identified, you can request another Copilot review.

Review details

Suppressed comments (1)

scripts/vllm_multinode/run_multinode.slurm:158

  • The barrier port is configurable via CONTAINER_BARRIER_PORT (and serve_colocated.sh honors it), but this fuser call always kills 2223. If the port is overridden, stale listeners on the configured port won’t be cleared and the barrier can hang/fail. Use the same ${CONTAINER_BARRIER_PORT:-2223} default here.
fuser -k 2223/tcp 2>/dev/null || true
  • Files reviewed: 15/16 changed files
  • Comments generated: 1
  • Review effort level: Lite

Comment thread scripts/vllm_multinode/run_multinode.slurm
Job 240624 spent its entire 7200s allocation in the pre-warm added by the
previous commit. Both nodes printed "[prewarm] reading ..." and neither ever
printed "[prewarm] done", so vllm serve was never launched -- the run timed
out having tested nothing.

The flaw: it reads the WHOLE checkpoint on EVERY node, while PP2xTP8 means a
node only ever loads its own shard. On a 1453 GiB checkpoint that is roughly
4x the necessary I/O against a single NFS export, with both nodes competing
for it. It manufactured exactly the contention it was meant to relieve.

Default it to off. It stays available, because the per-node duration is a
useful storage probe -- a large spread between nodes means the storage path is
the problem and no timeout will fix it -- but it is now opt-in via
PREWARM_CHECKPOINT=1 and bounded by PREWARM_TIMEOUT_SECONDS (default 900). On
expiry it warns that the filesystem could not deliver the checkpoint in the
window and proceeds, rather than burning the whole allocation.

Note this run did not exercise --cpu-distributed-timeout-seconds: vLLM never
reached the startup barrier, so the absence of the 1800000ms gloo timeout in
the log says nothing either way. That fix is still untested.

Verified: default unset is silent; PREWARM_CHECKPOINT=1 runs and reports its
duration; a stalled read hits the budget, warns, and exits 0 so startup
continues.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings September 3, 2026 18:19
Both GLM-5.1 cards passed at 1P/1D with each pool running TP8 and one DP rank:
cluster.sh exports TP_SIZE=GPUS_PER_NODE, and the wideEP path read TP_SIZE.
When the wideEP layout moved to its own knob, EP_TP_SIZE (default 1), these
cards silently became TP1 x 8 DP ranks. Every GPU then holds a full copy of
the non-expert weights, and decode crashes in MoRI EP setup with the GPUs
nearly full.

EP_TP_SIZE=8 on the cards restores the layout they passed with. The serve
argv matches the earlier one: -tp 8 is spelled --tensor-parallel-size 8, and
the pod-host list and --moriio-dp-size that EP_TP_SIZE>1 adds are one host
and 1 at 1P/1D, which is what the earlier path addressed. The recipe is
unchanged.
The check reads each card's own env_vars, so a card that loses EP_TP_SIZE
fails here: TP8, one DP rank per pool, and the single peer host in the
pod-host list, for both roles. It fails on the cards as they were before.
Copilot AI review requested due to automatic review settings September 29, 2026 16:06

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Cemberk and others added 2 commits September 29, 2026 22:24
pyt_vllm_disagg_deepep_deepseek-v3 still failed after 52ab5e1 gave it the
recipe's block size: decode logged "Memory access fault by GPU" on every
GPU while capturing its first decode cudagraph, right after loading AITER's
MLA kernel (mla_a8w8_qh128...), and the workers died.

The DeepSeek recipe turns AITER MLA off (VLLM_ROCM_USE_AITER_MLA=0), for
that reason, and asks for PIECEWISE decode cudagraphs. The deepep setup_env
exported VLLM_ROCM_USE_AITER_MLA=1 over it, and the decode launch hardcoded
FULL_DECODE_ONLY, so decode ran the ROCM_AITER_MLA backend in full-graph
mode. The moriio connector honours the recipe, picks TRITON_MLA with
PIECEWISE on the same model, and passes. The same lapse as 52ab5e1 and
4a28f2c.

setup_env now takes the AITER knobs the recipes set from the environment,
and decode takes DECODE_CUDAGRAPH_MODE (then VLLM_CUDAGRAPH_MODE) and
CUDAGRAPH_CAPTURE_SIZES, each defaulting to the value that was hardcoded,
so a recipe that sets none of them is unchanged. This also covers
pyt_vllm_disagg_nixl_deepseek-v3, which runs the same deepep path.

The dry run now prints the server's AITER env after the argv block, so
argv_assert can check it. argv_assert: 143 passed, 0 failed; four of the
new checks fail against the previous rixl.sh.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pyt_vllm_disagg_mori_kimi-k3 failed prefill start-up with CUDA out of
memory on every GPU while allocating the KV cache. Each GPU had 172.6 GiB
free before loading (the 16 GiB MoRI heap already taken), the TP2 weights
took 148.56 GiB, and KV_CACHE_MEMORY_BYTES=40e9 then asked for 37.25 GiB of
the ~24 GiB left. The recipe comment's budget assumed 137.5 GiB of weights,
which even then left no headroom.

16e9 (~1.13M tokens) leaves ~9 GiB for activations, decode cudagraphs and
RCCL, and still covers --max-model-len 1000000. Kimi-K3-MXFP4 passes NIAH
up to 280K tokens at the same TP2 x DP8 layout with 8e9. Single requests
beyond roughly 250K tokens may no longer fit; the card's NIAH sweep stays
within that. The MI355X recipe keeps 40e9.

argv_assert: 145 passed, 0 failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 29, 2026 22:29

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Cemberk and others added 2 commits September 30, 2026 00:14
…that failed

parse_to_csv.py looked for a result's [RUNNING] header in the text since the
previous result and took the first one there. When a cell stalls it prints
no result block, so the next cell's text holds both headers: on
pyt_vllm_disagg_deepep_deepseek-v3 the con=256 result (half its requests
failed) was filed as con=128, and the con=256 row vanished. It also never
looked before the first result, so the first measured cell of every sweep
was dropped: a passing DeepSeek-V3 MoRI run lost its con=8 row (512.27
tok/s).

Every sweep row was also written as SUCCESS. That run's con=512 cell failed
all 1024 requests and was reported as "SUCCESS 0.00 tok/s", and the build
passed.

The parser now takes the last header before each result, including the
first, records failed requests, and gives a [STALL] cell a zero-throughput
row. A cell that stalled, lost any request, or measured no throughput is a
FAILURE row, as agentic rows already are. The CONCURRENCY.csv keeps its
columns.

tests/parse_to_csv_assert.sh: 7 passed, 0 failed; 5 of them fail against the
previous parser.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rker OOM

With the KV cache fitting, pyt_vllm_disagg_mori_kimi-k3 got its prefill
master up and then lost decode: one decode worker ran out of GPU memory on
its first cudagraph capture (PIECEWISE, the largest size), when AITER's
fused-MoE asked for a 3 GiB workspace with ~9 GiB left after the 148.56
GiB of weights and 14.9 GiB of KV cache. Decode captured sizes up to 256
while the recipe runs --max-num-seqs 8, so every size above 8 was memory it
could never use. The recipe now sets CUDAGRAPH_CAPTURE_SIZES="1 2 4 8".

The OOM did not stop the job: the other workers waited, the engine never
printed an initialization failure, and the job sat ~55 min until vLLM's own
3600s engine-ready timeout. The start-up watch now treats a worker's
torch.OutOfMemoryError as fatal, as it already does an NCCL or HIP failure.

argv_assert: 149 passed, 0 failed; two of the new checks fail against the
previous recipe and launcher.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 30, 2026 01:46

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Cemberk and others added 2 commits September 30, 2026 01:49
Reverts the recipe side of 8370e44 (KV_CACHE_MEMORY_BYTES 40e9 -> 16e9) and
b3c5702 (CUDAGRAPH_CAPTURE_SIZES "1 2 4 8"). Both made the MI300X Kimi-K3
disagg card fit in memory by giving up what the recipe asks for: the
2.84M-token KV cache for single requests beyond ~600K tokens, and its
decode cudagraph sizes. The recipe is restored as it was; the out-of-memory
it hits on MI300X is to be fixed, not traded away.

What the two failed runs measured, for whoever takes that on: 148.56 GiB of
weights per GPU at TP2 (the recipe comment budgets 137.5 GiB), 172.6 GiB free
before loading with the 16 GiB MoRI heap. The same image also logs "This
AITER build ignores the SiTUv2 activation on the packed-int4 MoE path and
would silently compute SiLU" (needs ROCm/aiter#4471).

Kept from b3c5702: the start-up watch treats a worker's
torch.OutOfMemoryError as fatal. argv_assert: 145 passed, 0 failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
pyt_vllm_disagg_mori_kimi-k3-mxfp4 came up on all four nodes, prefill and
decode, and then failed every NIAH request with HTTP 500: the router's
prefill call got 404 "The model `<path>` does not exist." The recipe (the
PR 241-validated one) starts vLLM with --served-model-name kimi-k3, while
benchmark_niah.sh requested MODEL_PATH, on the assumption that disagg
recipes never set a served name.

vllm_disagg.sh now resolves SERVED_MODEL_NAME from the recipe's own flags
(--served-model-name when present, else MODEL_PATH, vLLM's default), and
NIAH requests it unless NIAH_MODEL is set. It reads the recipe rather than
the router's /v1/models, which 503s under MoRIIO service discovery. No
recipe changes; every recipe without --served-model-name resolves to
MODEL_PATH exactly as before.

The dry run now also prints SERVED_MODEL_NAME. argv_assert: 149 passed, 0
failed; the four new checks fail against the previous launcher.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 30, 2026 13:08

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

pyt_vllm_kimi_k3_mi300x, pyt_vllm_kimi_k3_mi355x and
vllm_disagg_inference.kimik3 were one stack three times. vLLM (862bfd8, the
head of the branch the disagg file named), MoRI (v1.2.2) and AITER were
identical in all three; they differed in the GPU arch, the router pin
(the disagg file's was 27 commits newer, and colocated cards do not use the
router), and NIXL on or off.

Each also built on the ROCm vLLM CI base, installed AITER 0.1.19, then
deleted it and copied aiter, aiter_meta and flydsl out of a second, "proven"
image. Both proven images (amdsiloai/vllm:kimi-k3-mi325x-release-v2 and
vllm/vllm-openai-rocm:kimi-k3) carry the same layer, digest e2b3951e36ca:
a checkout of carlushuang/aiter-k3 branch k3-for-amd at ROCm/aiter
68e42f5f (0.1.17.dev395), with the Kimi-K3 tuned MoE configs and
aiter.ops.triton.conv, and flydsl 0.2.4. (The old header's "the donor's
flydsl is 0.2.2" was wrong.)

docker/vllm_kimi_k3.ubuntu.amd.Dockerfile builds that AITER commit from
source, the same way the GLM-5.1 image builds its AITER, so there is one
base and no second image. It builds for MAD_SYSTEM_GPU_ARCHITECTURE
(gfx942 or gfx950, no default: a wrong-arch image fails at runtime, so the
build refuses to guess), the build arg madengine already passes; the arch
reaches each step as K3_GFX_ARCH so madengine's Dockerfile arch check does
not read a fixed arch. The router takes the disagg pin by SHA (82dc981),
rocSHMEM its pin, and WITH_NIXL defaults on so the image serves every
connector. Checks read package metadata only: importing aiter probes the
GPU, and the build agent has none.

A prefix no other Dockerfile shares, so madengine's `<prefix>.*` lookup
finds only this file even without its exact-match fix. All seven Kimi-K3
cards point at it; every card in scripts/*/models.json resolves to exactly
one Dockerfile. The MI355X bring-up caveats move into its STATUS.

argv_assert: 149 passed; parse_to_csv_assert: 7 passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 30, 2026 14:14

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Cemberk and others added 6 commits September 30, 2026 14:24
The Kimi-K3 image builds for MAD_SYSTEM_GPU_ARCHITECTURE and refuses to
guess; madengine cannot detect it on a build host without a GPU. The m2m
profile now states gfx942 as docker_build_arg, so a build with no probed or
given arch still gets the cluster's. An arch the CI probes or is given
overrides it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sets

scripts/common/clusters/m2m.json held partition amd-rccl, 8 GPUs per node
and exclusive, which are madengine's own SLURM presets and are also set in
cluster.sh. Nothing in MAD read it; only the MAD CI resolver did. The build
arch added to it in 08c45d3 moves to the CI's own target table (rocAutomation
NODE_PROFILES: oci-64 -> gfx942), which is where knowledge of CI targets
belongs.

Removed: clusters/m2m.json, clusters/README.md and the .gitignore exception
for them. cluster.sh and scripts/common/README.md now point at madengine's
presets and --additional-context instead of the profile. The allocation a
run gets is unchanged: the CI's path-equivalence check resolves all 46
multinode cards to --partition=amd-rccl --gpus-per-node=8 --exclusive, the
values the profile supplied.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
run_xPyD_models.slurm ended with the container-cleanup srun, so its exit status,
and the SLURM job's, was the cleanup's. A run whose prefill died at start-up
(pyt_sglang_disagg_mori_io_agentic_qwen3-32b, GPUs already occupied on a node)
ended COMPLETED with no results, and madengine collected "0 successful, 0
failed" and reported success.

The job now fails when the main srun exits non-zero (a node's container exits
non-zero when its server or benchmark fails; the passing runs checked exited 0 on
every node) or when no /run_logs/$SLURM_JOB_ID/perf.csv was written, as the vLLM
launchers already require. SKIP_BENCHMARK=1 serves without benchmarking and is
not required to produce results.

Checked with a stubbed srun: clean run 0; node failure 1; no perf.csv 1;
SKIP_BENCHMARK=1 without perf.csv 0; SKIP_BENCHMARK=1 with a node failure 2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
parse_to_csv.py took each result block's concurrency from the header before it
and wrote every row as SUCCESS. On pyt_sglang_disagg_mori_io_qwen3-32b the
con=256 cell lost half its requests to connection errors (256 of 512
succeeded) and was recorded SUCCESS at 6390 tok/s, and the con=512 cell, whose
warmup got "Bad Gateway" and printed no result, had no row at all; the build
passed.

Each RUNNING line now owns the text up to the next one. A cell that printed no
result is a FAILURE row at its own concurrency; sglang.bench_serving prints no
failed count, so a cell whose successful requests fall short of its prompts is
FAILURE; zero throughput is FAILURE. Clean runs are unchanged: re-parsing the
archived logs of two passing runs gives the same 7 SUCCESS rows and numbers.

tests/parse_to_csv_assert.sh (new): 7 passed; 4 fail against the previous parser.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d, and publish long_context results

Three gaps in how vLLM disagg results become perf.csv rows, each of which let a
broken run pass or a working one fail:

- A sweep cell whose benchmark crashed without the harness's [STALL] line (for
  example a warmup that got an error from a dead server) printed no result and
  got no row. parse_to_csv.py now reads the log cell by cell, as the SGLang
  parser does: each [RUNNING] line owns the text up to the next one, and a cell
  with no result is a FAILURE row.
- NIAH: a context length whose every request errored prints NO-RESULT and got no
  row, and a length that lost some requests ("[k timeout/err excluded]") was
  SUCCESS on the survivors' mean. Both are now FAILURE rows; a low or truncated
  retrieval score is still a measurement and stays SUCCESS. madengine counts
  FAILURE rows as failed runs and exits non-zero, so an all-errored NIAH run
  still fails, now with a row per length saying which.
- benchmark_long_context.sh never wrote perf.csv, so a long_context run always
  ended "no perf CSV" and failed. It now calls the parser as the sweep does,
  and the parser reads its "[RUNNING] isl=... con=..." header form.

Re-parsing archived logs: three passing sweeps and three passing NIAH runs are
unchanged; the all-NO-RESULT Kimi-K3-MXFP4 run gives four FAILURE rows. The
long_context script, run end to end with a stubbed vllm, writes a SUCCESS row
for a completed cell and a FAILURE row for a timed-out one.
tests/parse_to_csv_assert.sh: 16 passed; the new cases fail against the
previous parser.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…extending MAD

MAD had no docs folder; what a reader needed was spread over the root README and
READMEs nested under scripts/ and benchmark/. docs/ now holds it in one place, in
depth, in the order someone new needs it:

- README.md: what MAD is, a learning path, a glossary, and an index (with the
  madengine docs it builds on)
- getting-started.md, adding-a-model.md: running models; every card field,
  Dockerfile matching, MAD_SYSTEM_GPU_ARCHITECTURE
- multinode-overview.md, multinode-running.md: disaggregated and colocated
  inference, launchers, connectors, topology, the architecture diagrams; running
  through madengine, sbatch and salloc, weights, logs, failure modes
- configuration.md: every configuration layer, what wins, and where to change what
- benchmarks-and-results.md: each benchmark, perf.csv rows and what makes a run pass
- vllm-disagg.md, sglang-disagg.md, kimi-k3.md: full references

Where a README disagrees with the code, the pages follow the code. The existing
READMEs stay. The two this PR added, scripts/common/README.md and
benchmark/kimi_k3/mi300x/README.md, now point into docs/, and the root README links
to docs/README.md.

benchmark/kimi_k3/mi300x/slurm-config.{colocated,disagg}.json are removed, with
their .gitignore exception. They were fill-in templates that restated values with a
source elsewhere: the partition, GPUs per node, exclusivity and output directory
are madengine's SLURM presets, the node count is each card's, and MODEL_DIR and
LOG_PATH default in scripts/common/cluster.sh. The docs pass only what differs, as
--additional-context, and say where every other value comes from.

The Kimi-K3 Dockerfile's build note no longer says the context must be the repo
root: the image copies nothing from the build context.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 30, 2026 15:36

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

… have it

Both Kimi-K3 image builds failed at the AITER step with "fatal: reference is not
a tree: 68e42f5f". The release images built AITER from branch k3-for-amd of a
fork that is not public; the commit is in ROCm/aiter's fork network (GitHub
serves it by SHA and its API finds it) but on none of ROCm/aiter's branches, so
git clone never downloads it.

The step now fetches that exact commit by SHA, with --tags for the release tags
its version is derived from, and checks it out. Checked outside docker: HEAD
68e42f5f4, the composable_kernel submodule, the six kimik3 tuned configs and
aiter/ops/triton/conv are all present. The version string is now derived from
v0.1.18 (tags added upstream since the release images were built) instead of
0.1.17.dev395; the code is the same commit, and the image's check requires only
that the version names 68e42f5f.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 30, 2026 16:05

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants